Papers with Optical Character Recognition

22 papers
Optical Character Recognition for the International Phonetic Alphabet (2026.eacl-short)

Copied to clipboard

Challenge: Grammar books are increasingly used as additional reference resources for low-resource languages . a significant portion of these documents come from scans and require an OCR tool .
Approach: They compare two neural OCR frameworks and a large vision-language model with a synthetic dataset based on Wiktionary to study the International Phonetic Alphabet (IPA).
Outcome: The proposed model improves on the International Phonetic Alphabet (IPA) character set.
InkSight: Towards AI-Aided Historical Manuscript Analysis (2026.eacl-demo)

Copied to clipboard

Challenge: Large-scale scientific research on medieval Arabic manuscripts remains challenging due to the need for advanced paleographic and linguistic training and the lack of assisting software.
Approach: They propose an end-to-end Arabic manuscript analysis tool for manuscript-based analytics and research hypothesis testing.
Outcome: The proposed tool overcomes the limitations of existing tools and can be used in large-scale scientific research.
Empirical Error Modeling Improves Robustness of Noisy Neural Sequence Labeling (2021.findings-acl)

Copied to clipboard

Challenge: Standard sequence labeling systems fail when processing noisy user-generated text or consuming the output of an OCR process.
Approach: They propose an empirical error generation approach that employs a sequence-to-sequence model trained to perform translation from error-free to erroneous text.
Outcome: The proposed method outperforms baseline noise generation and error correction techniques on the erroneous sequence labeling data sets.
M3T: A New Benchmark Dataset for Multi-Modal Document-Level Machine Translation (2024.naacl-short)

Copied to clipboard

Challenge: Document translation is a challenge for machine translation systems that focus on textual content at the sentence level, ignoring global context and visual layout structure.
Approach: They propose a benchmark dataset to evaluate document-level NMT systems . they use visual cues to preserve reading order and contiguous blocks of text .
Outcome: The proposed benchmarks assess document-level NMT systems on the comprehensive task of translating semi-structured documents.
Language, OCR, Form Independent (LOFI) pipeline for Industrial Document Information Extraction (2024.emnlp-industry)

Copied to clipboard

Challenge: Existing models for low-resource language (LRL) documents are limited in semantic entity extraction, and there are limitations in SER from word level results.
Approach: They propose a pipeline for Document Information Extraction (DIE) in low-resource language (LRL) business documents that solves language, Optical Character Recognition (OCR), and form dependencies through flexible model architecture, a token-level box split algorithm, and the SPADE decoder.
Outcome: Experiments on Korean and Japanese documents show that the pipeline performs well in the Semantic Entity Recognition task without pre-training.
Books of Hours. the First Liturgical Data Set for Text Segmentation. (2020.lrec-1)

Copied to clipboard

Challenge: Until now, the book of hours has been scarcely studied because of its manuscript nature, its length and its complex content.
Approach: They propose to use Handwritten Text Recognition to generate a corpus of Latin transcriptions of 300 books of hours generated by OCR for handwritten and not printed texts.
Outcome: The proposed structure and state-of-the-art methods are compared with existing methods and are based on the results of a systematic evaluation of two books of hours.
Correction of OCR Word Segmentation Errors in Articles from the ACL Collection through Neural Machine Translation Methods (L18-1)

Copied to clipboard

Challenge: Optical Character Recognition (OCR) can produce a range of errors depending on the quality of the original document.
Approach: They applied a sequence-to-sequence machine translation system to correct word-single-word OCR errors in scientific texts from the ACL collection with an estimated precision and recall above 0.95 on test data.
Outcome: The proposed system corrects word-segmentation OCR errors with an estimated precision and recall above 0.95 on test data.
Vision Language Model Helps Private Information De-Identification in Vision Data (2025.findings-acl)

Copied to clipboard

Challenge: Visual Language Models (VLMs) have gained popularity due to their ability to solve imagerelated tasks.
Approach: They propose a framework to enhance privacy awareness of visual language models . they use a specialized instruction-tuning dataset and a tailored training methodology .
Outcome: The proposed framework outperforms existing approaches in handling private information.
LOCR: Location-Guided Transformer for Optical Character Recognition (2024.findings-emnlp)

Copied to clipboard

Challenge: Academic documents are packed with texts, equations, tables, and figures, posing challenges for accurate OCR results.
Approach: They propose a model that integrates location guiding into the transformer architecture during autoregression.
Outcome: The proposed model outperforms existing methods on an original large-scale dataset comprising 53M text-location pairs from 89K academic document pages.
How Much Data Do You Need? About the Creation of a Ground Truth for Black Letter and the Effectiveness of Neural OCR (2020.lrec-1)

Copied to clipboard

Challenge: Recent advances in Optical Character Recognition and Handwritten Text Recognition have led to more accurate text recognition of historical documents.
Approach: They propose to build a ground truth for a German-language newspaper published in black letter . they also evaluate the performance of different OCR engines and estimate how much data is needed to achieve high-quality OCR results.
Outcome: The proposed model can recognise black letter text and performs well on data they have not seen during training.
MCS-Bench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in Chinese Classical Studies (2025.acl-long)

Copied to clipboard

Challenge: Multimodal Large Language Models (MLLMs) have advanced visual and language understanding, but their potential in Chinese Classical Studies (CCS) remains underexplored due to the lack of specialized benchmarks.
Approach: They propose to develop a multimodal benchmark specifically designed for Chinese Classical Studies across multiple subdomains to bridge this gap.
Outcome: The proposed benchmark spans seven core subdomains with a total of 45 meticulously designed tasks.
UMRSpell: Unifying the Detection and Correction Parts of Pre-trained Models towards Chinese Missing, Redundant, and Spelling Correction (2023.acl-long)

Copied to clipboard

Challenge: Chinese Spelling Correction (CSC) is a task of detecting and correcting misspelled charac- ters in Chinese texts.
Approach: They propose a model to learn detection and correction parts together from a multi-task learning perspective.
Outcome: The proposed model can learn detection and correction parts together from a multi-task learning perspective.
Page Stream Segmentation with Convolutional Neural Nets Combining Textual and Visual Features (L18-1)

Copied to clipboard

Challenge: (retro-)digitizing paper-based files is a major undertaking for private and public archives and an important task in electronic mailroom applications.
Approach: They propose to use convolutional neural networks to combine image and text features to achieve optimal document separation.
Outcome: The proposed approach achieves an accuracy of 93 % and is considered a state-of-the-art for this task.
Evaluating Transformers for OCR Post-Correction in Early Modern Dutch Theatre (2025.coling-main)

Copied to clipboard

Challenge: a new study examines the effectiveness of two types of transformer models for OCR post-correction in early modern Dutch plays.
Approach: They propose to use large generative models and sequence-to-sequence models for OCR post-correction in early modern Dutch plays.
Outcome: The proposed model outperforms generative models on the OCR post-correction task . the model outpersforms the model with the lowest error rate on the historical English dataset .
Resilience of Large Language Models for Noisy Instructions (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) are powerful tools for interpreting human commands and generating text.
Approach: They examine the resilience of large language models against five common types of disruptions including ASR, OCR, grammatical errors, typographical errors and distractive content.
Outcome: The models show resistance to noise, but their performance suffers . authors evaluated the models against five common types of disruptions based on their results .
Unveiling the Power of Integration: Block Diagram Summarization through Local-Global Fusion (2024.findings-acl)

Copied to clipboard

Challenge: Document Artificial Intelligence (Document AI) is gaining momentum across industries for streamlining document processes, enhancing efficiency, and extracting insights from unstructured data.
Approach: They propose a fusion framework that summarizes block diagrams by integrating local and global information, catering to both English and Korean languages.
Outcome: The proposed framework surpasses all previous methods and models for block diagram summarization on a dataset of BD-EnKo in English and Korean.
StruNRAG: Evaluation of OCR-Induced Structural Noise on RAG Robustness (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluations of RAG systems ignore structural noise, authors say . complex layouts can cause OCR failures and disrupt semantic flow of text . advanced LLMs demonstrate robustness against local noise, but struggle to maintain reasoning capabilities under severe structural disruption that fragments global context.
Approach: They propose a benchmark to evaluate RAG robustness against OCR-induced structural perturbations.
Outcome: The proposed benchmark systematically injects three categories of real-world structural noise into a bilingual dataset of 2,132 question-answer pairs . results show that advanced LLMs demonstrate robustness against local noise, but struggle to maintain reasoning capabilities under severe structural disruption .
A Dual-View Analysis of Multiple Languages in Colonial Newspapers (2026.findings-acl)

Copied to clipboard

Challenge: Historical newspapers from the colonial period offer valuable evidence of how racializing language evolved over time.
Approach: They propose a contextual question answering and visual question answering task from colonial newspapers . they propose linguistic training for temporal word embedding with a compass to study racialization .
Outcome: The proposed tasks are limited for low-resource tasks, the authors show . the authors compare the results of two QA pairs from colonial newspapers to a compass .
PEaCE: A Chemistry-Oriented Dataset for Optical Character Recognition on Scientific Documents (2024.lrec-main)

Copied to clipboard

Challenge: Existing open-source OCR models focus on scientific texts or generic printed English . Nougat is unable to parse tables in PubMed articles .
Approach: They propose to train OCR models for scientific or generic printed English . Nougat is a popular tool for parsing academic documents, but unable to parse PubMed tables .
Outcome: The proposed models perform better when trained on real-world records than those trained on synthetic records.
KITAB-Bench: A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding (2025.findings-acl)

Copied to clipboard

Challenge: Optical Character Recognition (OCR) is a key component of document processing . Arabic text recognition has complex typographic and calligraphic features .
Approach: They propose a comprehensive Arabic OCR benchmark that fills the gaps in evaluation systems.
Outcome: The proposed benchmark outperforms existing models in Arabic by 60% in the character error rate . the best model achieves only 65% accuracy in PDF-to-Markdown conversion .
Improving MLLM’s Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency (2025.findings-acl)

Copied to clipboard

Challenge: Multimodal Large Language Models (MLLMs) have shown strong performance in document image tasks, especially Optical Character Recognition (OCR). However, they struggle with Document Image Machine Translation (DIMT), which requires handling both cross-modal and cross-lingual challenges.
Approach: They propose a novel fine-tuning paradigm that allows the model to generate OCR text before producing translation text, which allows it to leverage its strong monolingual OCR ability while learning to translate text across languages.
Outcome: The proposed model can leverage its strong monolingual OCR ability while learning to translate text across languages.
There’s Something New about the Italian Parliament: The IPSA Corpus (2024.lrec-main)

Copied to clipboard

Challenge: despite their potential, the Italian parliamentary documents remain unexplored and inaccessible in their original paper-based form.
Approach: They propose to transform Italian parliamentary documents into a structured corpus . the corpus includes speeches, reports of Standing Committees, and law proposals .
Outcome: The proposed dataset spans 175 years of Italian history spanning from the issuing of the Statuto Albertino in 1848, up to the present day .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations